BioData Mining
○ Springer Science and Business Media LLC
Preprints posted in the last 30 days, ranked by how well they match BioData Mining's content profile, based on 22 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Shenoy, A.; Zekarias, A.; Viklund, A.; Mitchell, J.; Barrett, J.; Sandberg, L.; Meldau, E.-L.; Taavola-Gustafsson, H.
Show abstract
Background Large Language Models (LLMs) are increasingly explored for pharmacovigilance tasks, including information extraction, case documentation, and single-case causality assessment. However, their ability to support causality assessment at the case series level -- a complex, time-intensive task requiring clinical reasoning across multiple reports -- remains unexplored. Objective To investigate how a large-scale general-purpose LLM can support pharmacovigilance professionals in assessing causality in a case series, and to explore how prompt design influences the quality of the model's reasoning. Methods GPT-4o was used to assess causality for five drug - adverse event combinations, using an adaptation of the Bradford Hill viewpoints for case series assessment. The combinations represented varying drugs and vaccines, adverse events, and case series sizes (5-402 reports). One combination served as a negative control. Structured prompts were iteratively developed and refined using one combination, then applied to all combinations. LLM-generated assessments for each viewpoint were qualitatively evaluated by human annotators for accuracy (precision), and the LLM's coverage of key aspects from the original signal text was assessed for one combination (recall). Results Across all five combinations, annotators agreed with 79-92% of the LLM's output sentences. Full disagreement was consistently low (3-7%), with errors typically involving misinterpretation of complex report details rather than outright fabrication. Prompt design substantially influenced output quality; providing Bradford Hill viewpoint descriptions, including case series data, and adding explicit anti-hallucination instructions improved specificity and grounding. For the recall assessment, 15 of 23 key segments from the original signal text were reflected in the LLM output. The overall summary assessments demonstrated balanced reasoning, correctly distinguishing between positive safety signals and the negative control, and provided a coherent synthesis suitable as a starting point for human assessors. Conclusions LLMs have the potential to generate contextually nuanced and largely accurate preliminary causality assessments of case series aligned with the Bradford Hill viewpoints, with a low but non-zero hallucination rate. These findings support LLMs as a tool to augment, not replace, expert judgment in signal assessment. Future work should address larger and more diverse signal sets, improved evaluation frameworks for generative output, and the integration of pre-computed summary statistics to reduce errors.
Kranz, A.-C.; Schneider, J.; Gassner, C.; Bublitz, M.
Show abstract
Blood group antigens, defined by epitopes on the erythrocyte surface, are central to transfusion safety and maternal-fetal compatibility. While the genetic basis of many clinically relevant blood group antigens is well established, which structural and biophysical parameters determine whether a single-nucleotide variant gives rise to an antigenic phenotype remains unclear. Here, we integrate structural, biophysical, and evolutionary analyses to systematically evaluate features associated with single amino acid substitutions across 24 human protein-based blood group systems. We analyse 319 variants with curated phenotypic annotations alongside 481 control variants, identifying key determinants of null and antigenic phenotypes. Null variants are characterized by high evolutionary conservation, burial within the protein core, loss of hydrophobicity, increased polarity, and a propensity for arginine substitutions. Antigenic variants are also enriched in arginine; however, in contrast to null variants, they tend to occur at less conserved, more solvent-accessible, and structurally flexible sites. Supervised machine learning models trained on structural and biophysical descriptors were applied to distinguish (i) null and (ii) antigenic variants from controls, achieving balanced accuracies of 0.82 and 0.63, respectively. Feature importance analysis identified predicted pathogenicity, solvent accessibility, and evolutionary conservation as the most predictive determinants of null variants, whereas hydrophobicity, conservation, and flexibility dominated antigen prediction. This work establishes a framework linking molecular variation to blood group phenotypes and provides a foundation for predicting the impact of novel missense mutations in transfusion medicine and beyond.
Asumboya, W. A.; Agbenorhevi, P. K.; Adams, C. F.; Ayariga, D. A.; Adjadeh, T.; Adams Ziblim, S.; Kwofie, S. K.
Show abstract
BackgroundClinical complications are often predicted with separate sigmoid outputs, even when the target labels arise from related pathophysiological processes. This paper asks whether output-layer choice should reflect both predictive convenience and the biological structure assumed among complications. The central premise is that label-dependence mechanisms are explicit hypotheses about comorbidity, not generic modelling additions. MethodsOutput-head assumptions were compared across two clinically distinct multi-label prediction tasks. In Type 2 diabetes (T2D), six heads were evaluated for nephropathy, neuropathy, and retinopathy: independent baseline, linear additive, multiplicative, symmetric conditional random field (CRF), residual multilayer perceptron (MLP), and combined additive-multiplicative. In myocardial infarction (MI), four heads were evaluated for ventricular tachycardia, ventricular fibrillation, and atrioventricular block: independent baseline, linear additive, multiplicative, and symmetric CRF. All experiments used five training data fractions and seven independent seeds, with the same shared-backbone protocol within each disease setting. ResultsIn T2D, the symmetric CRF gave the most consistent improvement pattern, ranking highest at full data and at the two lowest data fractions while adding only three interaction parameters. At 20% training data, it was the only interaction head whose aggregate mean exceeded the independent baseline. The residual MLP, despite 123 interaction parameters, remained below the baseline across all T2D fractions. In MI, rankings changed across fractions: the multiplicative head led at 80% and 60%, the CRF led at 100% and 20%, and the baseline led at 40%. The combined additive-multiplicative head did not improve robustness in T2D and showed the largest negative baseline-relative deviations at lower fractions. ConclusionThe findings support a biology-guided view of output-layer design. A small constrained mechanism was most useful when its symmetry matched the shared microvascular structure of T2D, whereas the heterogeneous electrophysiology of MI produced no stable winner. Output-layer choice should therefore be reported and defended as an assumption about disease structure instead of a routine hyperparameter decision. Author summaryMany clinical prediction models treat complications as separate outcomes, even when clinicians know they often arise together. We studied whether the last layer of a model should reflect that biological knowledge. We compared several output heads across two disease settings: Type 2 diabetes, where nephropathy, neuropathy, and retinopathy share a common microvascular origin, and myocardial infarction, where electrical complications arise from a mixture of shared and location-specific mechanisms. We found that a small symmetric CRF head was most useful in the diabetes task, especially when training data were limited, while no single interaction head dominated in myocardial infarction. This suggests that modelling comorbidity is not only a technical choice; it is a statement about how disease processes relate to one another. Our results encourage researchers to report and justify output-layer design as part of the clinical modelling argument, rather than treating it as a routine hyperparameter.
Whalley, J. P.
Show abstract
BackgroundSingle-cell foundation models are increasingly used for perturbation prediction and gene network inference, but their learned gene representations are rarely audited directly. In natural language processing, geometric analyses of token embeddings have revealed anomalous "glitch tokens" associated with erratic model behaviour. Whether analogous representational anomalies exist in biological foundation models remains unknown. ResultsThis study introduces a weight-only geometric audit framework that scores genes by embedding norm, centroid distance, cosine similarity, and isolation to identify representational outliers. Applied to Geneformer, scGPT, and scFoundation, the analysis identifies hundreds of outliers in discrete-tokenisation models. Shared Geneformer-scGPT outliers are enriched for loss-of-function intolerance (OR=12.0) and disease association (OR=3.7), whereas scFoundations continuous value embeddings form a near-isotropic space with no detectable enrichment under the annotation panels tested. In Geneformer, geometric anomaly predicts perturbation sensitivity ({rho} = 0.725); the signal is supported by mask-in-place experiments, shows rank agreement in real PBMC cells, and correlates with Replogle perturb-seq effect sizes ({rho} = 0.645). Metric decomposition separates magnitude-driven outliers, enriched for highly expressed housekeeping genes, from isolation-driven outliers enriched for tissue-restricted genes. ConclusionsTokenisation strategy helps determine which genes are represented reliably. Embedding geometry provides a rapid, model-agnostic diagnostic that requires only an embedding matrix and can flag genes whose representations warrant caution before downstream use.
Stingl, J. C.; Molden, E.; Hole, K.; Wollman, B.; Viviani, R.
Show abstract
Background. Polypharmacy is an important source of phenoconversion caused by drug interactions potentially modulated by genetic variability. Aims. To develop a linear phenoconversion model for TDM data and provide quantitative estimates of drug-drug-gene interactions (DDGIs) in the pharmacogenetic phenotype groups of CYP2C19. Methods. Escitalopram TDM data in a large real-world sample (n=2,852) was analysed for phenoconversion of CYP2C19 activity. Co-medication was identified by reprocessing high-resolution mass-spectra (Orbitrap). We developed a statistical model to identify inhibition from co-medication in the CYP2C19 and in alternative elimination pathways. We extended the model to estimate the inhibition ensuing from individual co-medications, using a single model for all data to account for multiple co-medications and confounders simultaneously. A Bayesian approach allowed us to stabilize the fit and provide well-calibrated credibility intervals. Results. Reprocessing of TDM analyses identified 17 co-medications, which were shown to phenoconvert CYP2C19 activity proportionally to the activity in non-medicated phenotypes. Phenoconversion decreased the original CYP2C19 activity by about one third for a co-medication that corresponded to a 100% substrate of CYP2C19. The extent of CYP2C19 phenoconversion correlated strongly with the fractional contribution of CYP2C19 to the metabolism of the specific co-medication reported in the pharmacogenetic literature (R2=0.55) so long as the mechanism was competitive inhibition. Conclusion. We provide the statistical methodology to estimate phenoconversion from co-medication in TDM data and combine TDM and pharmacogenetic datasets in future studies aiming at establishing quantitative models of DDGIs.
Sarkar, P.; Sarkar, P.
Show abstract
Colorectal cancer (CRC) is challenging to track because its molecular changes are very complex as the disease progresses, creating significant challenges for robust biomarker discovery. In this study, we developed a machine learning framework by integrating monotonic progression and the StepMiner approach. We conducted external validation to identify reproducible, consistent transcriptomic biomarkers associated with CRC progression. Gene expression datasets were analyzed across four disease states from publicly available GEO: normal colon, adenoma, primary colorectal cancer, and metastasis. First, we identified genes with monotonic expression, then used the StepMiner approach to identify genes that act as switches between stages. A balanced 74-gene signature was used for machine-learning classification with a Random Forest. External validation showed strong performance in tissue-based datasets. However, tissue-derived signatures and plasma and blood-based datasets showed poor performance, highlighting biological differences between transcriptomic profiles. Cross-filtering between tissue-derived genes and blood expression datasets was performed, which resulted in the selection of 62 blood-compatible gene signatures. Leakage-free retraining on GSE164191 achieved a mean AUC of 0.868 with balanced precision. Functional enrichment analysis showed that these genes are highly active in cancer growth. Specifically, genes CBX3, S100A11, PDK4, NCOR1, and SOX4 demonstrated stable and reliable performance across the validation fold. Overall, our study presents a progression-aware transcriptomic framework for CRC biomarker discovery and demonstrates the importance of external validation. Additionally, we evaluate whether tissue-derived signatures can predict blood profiles. This proposed approach may help the future development of tissue-based diagnostics and minimally liquid-biopsy strategies for CRC. To ensure reproducibility, our proposed workflow was automated as a Nextflow pipeline. The tissue-derived model was deployed as an application utilizing Angular, ASP.NET Core, and Plumber (R).
Berardelli, S.; BRIERE, G.; Loire, B.; De Paoli, F.; Gazzo, A. M.; Limongelli, I.; Magni, P.; Zucca, S.; Baudot, A.
Show abstract
Motivation: Standardized phenotypic descriptions are essential for accurate diagnosis, yet clinicians and researchers face challenges in manually extracting and mapping phenotypes from scientific literature or patient clinical records to the Human Phenotype Ontology. Recent advances in deep learning offer new opportunities for automation. We developed PhenoXtract, a novel phenotype extraction approach that combines Large Language Models and Knowledge Graph embedding. PhenoXtract is a multistep pipeline that takes clinical descriptions as input, extracts candidate phenotype entities using large language models, and maps them to terms from an enriched version of the Human Phenotype Ontology, processed as a knowledge graph. Results: Evaluation against expert-curated ground-truth datasets show a recall of 0.70 and precision of 0.85 for PhenoXtract, demonstrating concordance with manually extracted phenotypes, with a computation time of 10-20 seconds for each text analyzed. Moreover, PhenoXtract surpasses rule-based and deep learning-based state-of-the-art tools in two out of the three ground-truth datasets evaluated. These results suggest that hybrid approaches combining Large Language Models and Knowledge Graph embeddings represent a promising direction for automated clinical phenotyping at scale.
Sundelin, H.; Jacobsson, B.; Ytterberg, K.; Sole-Navais, P.; Juodakis, J.
Show abstract
The leading cause of mortality and morbidity in children under the age of 5 is preterm birth. The timing of birth is influenced by both genetic and environmental factors, but the underlying mechanisms remain poorly understood, making its prediction difficult. In this study, we investigated the potential of using machine learning models to predict preterm birth based on genetic data from the Norwegian Mother, Father and Child Cohort Study (MoBa). We trained and evaluated several classification algorithms on individual-level genetic data from over 15,000 mothers and children. Our results indicate that the predictive capacity of maternal gestational duration-associated loci for preterm birth is limited, with the highest AUC values around 0.57. Additionally, incorporating more SNPs within the associated loci did not improve prediction performance. As expected, the contribution of the maternal genome to preterm birth prediction was found to be larger than that of the fetal genome. Overall, our findings suggest that while genetic testing provides some information about an individual's risk for preterm birth, further research incorporating additional factors is necessary to enhance predictability.
Chen, Z.; Wang, R.; Luo, Q.
Show abstract
Protein language models (pLMs) offer great potential for protein sequence analysis, yet the scarcity of labeled data often limits their effectiveness in fine-tuning. Data augmentation is a promising remedy, but systematic evaluation of augmentation strategies for protein sequences remains limited, and the conditions under which augmentation confers downstream benefits are not well understood. In this paper, we systematically investigate pLM-guided substitution-based augmentation across seven protein prediction tasks. We propose ProtAug, a framework that leverages encoder-based (ESM-2) and autoregressive (ProtGPT2) pLMs to generate augmented sequences with user-controlled variation levels. Our investigation focuses on four questions: (Q1) whether pLM-synthesized sequences preserve more original signals than simpler methods, (Q2) to what extent augmentation improves prediction performance, (Q3) how variation levels affect downstream accuracy across tasks and models, and (Q4) whether biological plausibility is a necessary condition for achieving improvement. Our experimental results show that: (1) ProtAug Esm generally preserves motifs and structural similarity better than simple substitution, often comparable to homology retrieval; (2) augmentation yields consistent but task-dependent improvements, with ProtAug Esm achieving the best or second-best performance in 5 out of 7 tasks at 10% variation; (3) low-to-moderate variation levels (2-30%) perform best overall, although high-variation augmentation can benefit certain structure-related tasks; (4) the necessity of biological plausibility is task- and variation-dependent--while semantic preservation correlates with performance at low-to-moderate variation levels, improved generalization at high variation levels suggests that regularization effects, rather than label preservation, can also drive performance gains.
Walker, A.
Show abstract
Genome mining is a powerful technique in natural product discovery, where biosynthetic gene clusters that are likely to produce novel or desirable natural products are identified through bioinformatic analysis. There are many more predicted biosynthetic gene clusters than can easily be experimentally characterized. Additional computational methods to prioritize biosynthetic gene clusters by the bioactivity, structural properties, or novelty of the product would make genome mining more efficient. Multiple machine learning/artificial intelligence models have been developed to predict product properties from biosynthetic gene cluster sequence, but they are limited by small quantities of training data. Model pretraining with unlabeled data is a powerful technique to develop models that can learn on a limited amount of labeled training data. Biosynthetic gene clusters are well suited to this strategy because there are many predicted clusters with only a small percentage being characterized. This paper reports BGC-MLM, a foundation model that is pretrained with a masked language task on predicted biosynthetic gene clusters and then fine-tuned for downstream applications including prediction of product structural class, bioactivity, chemical properties, counts of functional groups, and chemical fingerprint. Comparison to a model trained without pretraining shows that pretraining generally improves performance. BGC-MLM shows better or similar performance to existing specialized methods for these tasks, demonstrating its utility as a foundation model for natural product genome mining.
Pham, M.-D. N.; Phan, M.-T. T.; Tran, N.-T.; Vo, T.-S.; Le, H.-T.; Nguyen, T.-H. T.; Nguyen, Q.-H. V.; Ha, M.-T. T.; Le, T. M.; Hoang, D.-T. T.; Huynh, K.-T. N.; Nguyen, N. V.; Nguyen, C. C.; Bui, T. C.; Nguyen, X. T.; Le, S. V.; Tran, V. D.; Nguyen, M.-N. B.; Nguyen, T. V.; Nguyen, T.-A. T.; Hoang, B. P.; Nguyen, T. V.; Nguyen, T.-A. T.; Nguyen, T. T.; Duong, T. D.; Pham, C. H.; Luong, K.-O. T.; Dao, C. N.; Hoang, K. V.; Huynh, T.-T. T.; Nguyen, K. M.; Tran, S.-T. T.; Tran, H. T.; Nguyen, S. C.; Tran, T. D.; Nguyen, P. T. L.; Pham, T. V.; Pham, K. C.; Thai, M. D.; Do, T.-T. T.; Dao, H. T.; Va
Show abstract
ObjectiveTo develop and validate a cell-free DNA (cfDNA) fragmentomic classifier for the early prediction of spontaneous preterm birth (PTB) using routine first-trimester non-invasive prenatal testing (NIPT) data. MethodsA nested case-control study was conducted within a prospective multicenter Vietnamese cohort comprising 286 pregnancies, including 82 spontaneous PTB cases and 204 term controls. Maternal plasma cfDNA collected during routine first-trimester NIPT (median gestational age, 12 weeks) was sequenced to a depth of approximately 20 million reads per sample. Five fragmentomic feature categories including copy number alterations, end-motif composition, nucleosome distance, fragment length, and joint fragment-lengthxend-motif were evaluated for PTB prediction. Machine learning classifiers were developed in a training cohort (n = 228, 65 PTB vs 163TB) and tested in a validation cohort (n = 58, 17 PTB vs 41 TB). ResultsAmong the five fragmentomic feature classes evaluated, 4-mer end-motif (EM) profiles exhibited the most pronounced differences between PTB and term control samples. Consistent with these findings, the EM-based classifier demonstrated the highest discriminative performance in the validation cohort, achieving an AUC of 0.970 (95% CI, 0.912-1.000). At a specificity >90%, the model achieved a sensitivity of 94% (95% CI, 78-100%). ConclusionThese findings demonstrate that cfDNA EM signatures derived from routine first-trimester NIPT can accurately identify pregnancies at risk of spontaneous preterm birth, without additional blood collection or sequencing, thereby extending the clinical utility of existing prenatal screening infrastructure. KEY POINTSO_ST_ABSWhat is already known about this topic?C_ST_ABSO_LICurrent first-trimester prediction strategies based on maternal characteristics, cervical length, and biochemical markers have limited predictive accuracy, particularly in nulliparous women. C_LIO_LIExisting cfDNA-based approaches have shown only modest performance or require additional assays, limiting clinical applicability. C_LI What does this study add?O_LIExisting NIPT sequencing data can be repurposed (without additional blood sampling or sequencing) for accurate prediction of spontaneous preterm birth (AUC=0.970). C_LIO_LIA classifier employing 4-mer end-motif (EM) profiles achieved an AUC of 0.970. At a specificity >90%, the model achieved a sensitivity of 94%. C_LI
Konstorum, A.; Xing, J.; Aeron, S.; Kilmer, M.; Kleinstein, S.
Show abstract
Systems-level immune profiling data arising from longitudinal studies of vaccination or infection has an inherent multi-index array structure. While tensor decomposition of such datasets has gained popularity, choosing a rank and trial for a decomposition is not straightforward. We show that taking into account the experimental data model can inspire the development of new metrics to assess the quality of a Non-negative CANDECOMP/PARAFAC (NCPD) decomposition, and can thus be used to choose a rank and trial for the decomposition. Moreover, we show how framing the results via a dictionary learning framework can better enable interpretation of the components of the decomposition.
Bituin, R. C.; Bokani, A.
Show abstract
Systematic reviews in computational biology require screening large heterogeneous bibliographic sets, especially when topics span computational methods, cancer genomics and statistical modelling. This paper presents a reproducible semantic triage pipeline that combines SPECTER scientific-document embeddings, research-question similarity, proposal-summary similarity and domain keyword coverage to rank candidate studies for systematic review screening. The pipeline was evaluated on 2,231 Covidence records, including 120 final included studies (prevalence = 5.38%), against keyword-only, TF-IDF, BM25, MiniLM, PubMedBERT and SPECTER-only baselines. SPECTER-hybrid achieved the highest average precision (AP = 0.546), recovered 50% of included studies after screening 4.48% of records, and produced an 11.16-fold enrichment over prevalence. Ablation analysis showed that semantic-keyword combinations consistently outperformed single-signal variants. These findings suggest that citation-informed hybrid ranking can support literature triage while retaining human reviewers as final decision-makers.
Aselstyne, A.; Karthik, E. N.; El Azami, M.; Pogorelcnik, R.; Fournier, Q.; Chandar, S.
Show abstract
Motivation: Antimicrobial resistance (AMR) has been identified as a top global public health threat. Accurate AMR phenotype prediction from whole-genome sequencing data is an essential tool for accelerating clinical decision-making and mitigating resistance spread. Although many previous works have explored the use of tree-based machine learning (ML) models to predict resistance, the field lacks a systematic evaluation of the training pipeline across a variety of pathogenic species and antibiotics. Results: Using nine clinically relevant species-antibiotic combinations from the NCBI antimicrobial susceptibility testing database, we present a detailed analysis of the ML pipeline and identify key factors affecting model performance and evaluation. We begin by relabelling all isolates using current CLSI minimum inhibitory concentration breakpoints to resolve inconsistencies and increase available data, resulting in up to a 19% label swap and 56% data enlargement per species-antibiotic combination. We identify several key training parameters including k-mer length, which can increase classification F1 scores by over 20 points compared to commonly used k-values, feature matrix truncation, which can induce polynomial time reductions with limited performance reduction, and ML model class. By comparing 5-fold cross-validation with evaluation on an unseen clinical dataset, we show that random cross-validation splits--often criticized as overly optimistic--can act as a strong proxy for downstream clinical performance, yielding closer F1 scores than phylogeny-aware splits in all cases. We finally present an interpretability study which shows that over 95% of k-mers used by our models are associated with identifiable genomic features. Our results highlight the importance of feature design, evaluation protocol, and biological analysis in genomic AMR prediction, and support tree-based models as a robust and interpretable method.
Ahmed, M. O.; Amale, S. A.; Bhavsar, R. D.; Chopra, P.; Jaimes, A.; Kachhwah, A.; Kalotra, C. D.; Li, P.; Li, X.; Liao, Y.; Roy, R.; Senthilselvan, N.; Shao, Y.; Sharma, A. D.; Shrivatsan, A.; Xue, R.; You, Y.; Badkul, A.; Xie, L.; Oet, M.; Lee, K.; Sinitskiy, A.
Show abstract
Artificial Intelligence (AI) frameworks for automating scientific research have shown strong performance on benchmarks, but their capacity to routinely reproduce results from multiple real-life published studies remains largely untested. We evaluated five advanced AI research frameworks (Kosmos, K-Dense, ToolUniverse, BioAgents from bio.xyz, and the AI Scientist-v2 from Sakana AI) on three real-life tasks (including two recently published papers) spanning uncertainty quantification for molecular property predictions, machine learning on Therapeutic Data Commons benchmarks, and agent-based modeling. AI frameworks demonstrated genuine strengths: generating original hypotheses, competently executing routine data acquisition and coding tasks, providing statistical measures of confidence often absent from the original papers, and producing well-formatted final reports. At the same time, our experiments revealed that real-world scientific tasks remain considerably harder than current benchmarks suggest. No AI framework matched the scope or depth of the original studies, results varied across multiple runs of the same framework with the same prompt, and we documented cases of severe hallucinations in final reports, gaps in literature coverage, and overconfident conclusions. Verification of AI outputs required substantial domain expertise. While these three tasks are only partially representative of the broader scientific landscape, they offer a starting point for developing a more rigorous methodology for evaluation of AI performance than what is currently practiced. We conclude that AI frameworks are already valuable for prototyping research directions and stress-testing completed studies, and some of the limitations documented here appear largely tractable through infrastructure improvements and continued development.
Gulluoglu, H. S. A.; Baby, J.; Bagul, K. M.; Basangari, B. R.; Bathini, S. A.; Chalamalla, N. K. R.; Dcunha, J.; Gupta, O.; Huang, L.; Jiang, X.; Naidu, Y. R.; Sathishkumar, G.; Sehrawat, M.; Thota, S. L.; Thuvara, D.; Vanguri, M. B.; Yin, J.; Jugder, B.-E.; Lusky, I. E.; Li, J.; Sinitskiy, A.
Show abstract
Agentic artificial intelligence (AI) systems increasingly claim to automate scientific research, yet independent evaluations report persistent gaps between those claims and demonstrated capability. We tested frontier agentic AI systems on three practical problems: prediction of treatment non-response in immune-mediated inflammatory diseases, optical chemical structure recognition for literature mining, and prediction of drug-design-related properties from small datasets. Each problem was first assigned to autonomous frameworks and then reattempted as human-led, AI-assisted work. Autonomous runs failed in most cases, while human-led work produced reusable resources and modest but defensible performance, including new evidence for possible mechanisms of treatment resistance and a more practical benchmark for mining chemical structures from scientific papers. Property prediction was the single task on which one autonomous AI framework matched the human expert. We conclude that current frameworks can carry out engineering and analysis once a human expert leads the project, but cannot yet engineer a novel solution without oversight. The use of AI on real-life scientific problems remains an art rather than a routine technology.
Sautreuil, C.; Lesueur, C.; Pinto Cardoso, G.; Bruel, H.; Biran, V.; Muller, J.-B.; Duigou, A.-L.; Datin-Dorriere, V.; Verspyck, E.; Marguet, F.; Laquerriere, A.; Gressens, P.; Gonzalez, B.; Marret, S.
Show abstract
Prenatal alcohol exposure (PAE) is a major cause of neurodevelopmental disorders, yet most children are diagnosed late or misdiagnosed. Neuroplacentology suggest that placental factors released into maternal and/or umbilical cord blood contribute to fetal brain development. Consistently, a preclinical inter-organ transcriptomic database revealed that PAE disrupts the expression ratio of angiogenic and inflammatory factors suggesting an angio-inflammatory response. This study aimed i) to assay, by multiplex immunoassay, angiogenic and inflammatory factors in maternal and umbilical cord blood from alcohol-consuming women and ii) to perform a maternofetal analysis according to neonatal sex. Afterwards, dysregulated factors from mothers who gave birth to females or males were submitted to STRING and ShinyGO analyses. Results showed that PAE differently altered the distribution profiles of dysregulated angiogenic and inflammatory factors in maternal and umbilical cord blood. Moreover, sex-specific differences were observed, with 36% of dysregulated proteins specific to males, 48% to females, and 16% common to both. STRING analysis revealed robust functional protein-protein interactions linking together inflammatory and angiogenic clusters while the ShinyGO analysis identified enriched pathways related to vascular shear stress. These findings provide the first maternofetal analysis of combined angiogenic and inflammatory factors from alcohol-consuming mothers.
Bowman, M.; Bandopadhyay, R.; Singh, V.; Telpoukhovskaia, M.; Vander Velde, R.; Shaffer, S. M.; Trowbridge, J. J.; Bowman, R. L.
Show abstract
Single cell RNA-seq (scRNA) has provided unprecedented resolution into cellular and clonal heterogeneity. Computational approaches have enabled recovery of differentiation dynamics, yet current approaches do not evaluate discontinuous differentiation processes present in malignant leukemia. To address these gaps, we developed SupeRJump: a jump-drift-diffusion based supervised cell-fate model (https://github.com/namwob44/SupeRJump/). We deploy this approach in human bone marrow, murine aging hematopoiesis, and lentivirally barcoded mouse models of acute myeloid leukemia. Our framework introduces a semi-supervised pseudotime strategy to fit a jump-drift-diffusion model and batch correction for lineage fate predictions from absorbing Markov chains. We introduce metrics to quantify cell skewness toward particular lineages, transitions through intermediate progenitor states toward terminally differentiated states, and discontinuous transition dynamics. We use these metrics to identify cells preferentially biased for differentiation, their underlying transcriptional networks, and gene programs responsible for differentiation discontinuity.
Pestian, J. P.; Jacobson, D. A.; Pedapati, E. V.; Mendonca, E. A.; McMahon, B. H.; Ive, J.; Glauser, T. A.
Show abstract
The emotional content of suicide notes is typically examined using categorical coding, where each labeled passage is treated in isolation from its surrounding language. In contrast, dimensional models of psychopathology propose that affective content varies along continuous gradients. We evaluated this proposition directly. Excerpts from 884 annotated suicide notes were embedded in a semantic space defined solely by their linguistic properties, and we investigated whether human-assigned emotion labels changed smoothly across this space. They did: affective tone showed clear spatial autocorrelation (Moran's $I = 0.18$, $z = 19.68$, $p < 0.001$), an effect that replicated across three different encoders and remained after removing all within-note dependencies. Emotions occupied recognizable yet overlapping regions rather than forming distinct clusters and varied substantially in how tightly they were concentrated: love and hopelessness appeared with similar frequency, but love was far more localized ($z = 15.7$ versus $10.8$). Among all emotions, hopelessness was the most linguistically diffuse, implying that a single categorical label is capturing multiple, qualitatively different manifestations of suicidal distress.
Kadasova, N.; Martinat, D.; Spackova, A.; Hutarova Varekova, I.; Berka, K.
Show abstract
Significance Missense mutations can lead to pathological effects in human cells. Predictive methods that account for structural context, such as AlphaMissense, can provide pathogenicity scores. The accumulation of pathogenicity hotspots can reveal important structural features within individual proteins of protein families, such as GLUT transporters. Mapping pathogenicity scores onto the structure can thus provide a mechanistic explanation of the protein function necessary for its role in the cell. Abstract Non-synonymous amino acid substitutions (missense mutations) are common in the general population; some are causative of serious disease. Depending on their structural context, they can disrupt protein function, folding, or dynamics. Computational predictive methods developed in recent years, such as AlphaMissense, provide new insights into how missense mutations affect protein structure by predicting and mapping their pathogenicity across each amino acid in the human proteome. In this study, we identify recurring patterns of pathogenicity prediction across the GLUT family membrane transporters encoded by genes slc2a1-14. Within the GLUT transporter family, we observe higher pathogenicity profiles in the transmembrane domains, particularly in pore-lining and binding-site residues. Predicted missense pathogenicity is elevated throughout residues assigned to the central cavity, suggesting sensitivity of the transport pathway. Another finding shows higher pathogenicity in specific transmembrane helices of the protein, with the same pattern across all proteins. On the other hand, we observed lower pathogenicity values in some representatives of the GLUT family. These findings show that the pathogenicity of glucose transport within the GLUT family may be shaped by functional redundancy and physiological essentiality across GLUT groups.